Papers with low-level tuning

    1 papers
    HqeKV: Towards Hybrid Quantization and Eviction for KV Cache in Long-Context LLM Inference (2026.findings-acl)

    Copied to clipboard

    Challenge: autoregressive inference requires repeated computation across transformer layers.
    Approach: They propose a hybrid compression framework built on both quantization and eviction . they propose varying importance metric and flexible conversion policies to reduce memory overhead .
    Outcome: The proposed framework outperforms state-of-the-art methods under memory constraints.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations